FAU Erlangen-Nürnberg
Institute of Micro- and Nanostructure Research
regionprops / local-environment descriptors → trees, RF and GBM, validated with GroupKFold by image.Road map (≈90 min)
After today you can …
Self-study: notebooks/week06_tiny_mlp.ipynb
Left: linear regression on the hand-crafted \(1/\sqrt{d}\) feature achieves \(R^2\approx0.91\) (225 training samples). Right: a small MLP on raw grain size \(d\) only reaches \(R^2\approx0.65\) — with limited data, domain knowledge still wins over learned features.
regionprops; only \(w_j\) is learned.Hand-crafted features encode what the engineer knows. Learned features encode what the data contains.




\[\mathbf{z} = W\mathbf{x} + \mathbf{b}, \qquad \mathbf{a} = \sigma(\mathbf{z})\]
Step, sigmoid, tanh, ReLU, Leaky ReLU. ReLU preserves the gradient for positive inputs; sigmoid and tanh saturate.
| Task | Output activation | Loss |
|---|---|---|
| Regression | identity | MSE or MAE |
| Binary classification | sigmoid | binary cross-entropy |
| Multi-class | softmax | categorical cross-entropy |

\[\hat{R}(\theta) = \frac{1}{N}\sum_{i=1}^{N} \mathcal{L}\!\bigl(f_\theta(\mathbf{x}_i),\, y_i\bigr), \qquad \theta \leftarrow \theta - \eta\,\nabla_\theta \hat{R}\]
| Method | Passes needed | Cost for \(W\) params | Exact? |
|---|---|---|---|
| Finite differences | \(W+1\) forward | \(O(W^2)\) | no (\(\epsilon\) error) |
| Symbolic | — | expression swell | yes |
| Backprop (reverse AD) | 1 forward + 1 backward | \(O(W)\) | yes |

loss.backward() fills .grad.Network: 1 input → 2 ReLU hidden units → 1 linear output, squared loss.
\[ \begin{aligned} \mathbf{z}^{(1)} &= \mathbf{W}^{(1)}x + \mathbf{b}^{(1)}, &\quad \mathbf{a}^{(1)} &= \mathrm{ReLU}(\mathbf{z}^{(1)})\\ \hat{y} &= \mathbf{w}^{(2)\top}\mathbf{a}^{(1)} + b^{(2)}, &\quad L &= \tfrac12(\hat{y}-y)^2 \end{aligned} \]
Numbers: \(x = 1,\; y = 0.5\), \(\mathbf{W}^{(1)} = \begin{pmatrix}2\\-1\end{pmatrix}\), \(\mathbf{b}^{(1)} = \begin{pmatrix}-0.5\\0.5\end{pmatrix}\), \(\mathbf{w}^{(2)} = \begin{pmatrix}1\\2\end{pmatrix}\), \(b^{(2)} = 0\).
Forward values (black) and backward gradients (red) for the worked example. Each weight gradient = incoming error signal × stored forward input. The inactive ReLU unit (\(z_2^{(1)}<0\)) passes no gradient. Figure: img/make_figures.py.
torch.no_grad() for evaluation · model.eval() switches dropout / batch norm to inference mode.loss.backward() vs finite differences on the worked example and a random 2-layer net.NaN.clip_grad_norm_(model.parameters(), 1.0).Standard deviation of activations layer by layer in a 20-layer, width-256 network fed with unit-variance input. ReLU (left): small weights → activations vanish, \(\mathcal{N}(0,1)\) → explode, Xavier → slow decay, He → constant. tanh (right): large weights saturate at ±1, small weights vanish, Xavier/He stay usable. Figure: img/make_figures.py.
StandardScaler fit on the training split only (Week 4 leakage rule).model.eval()!| averages over | batch-size dependent | train ≠ eval | typical use | |
|---|---|---|---|---|
| BatchNorm | batch (per feature/channel) | yes | yes | CNNs, large batches |
| LayerNorm | features of one sample | no | no | transformers (W10), MLPs |
| GroupNorm | channel groups of one sample | no | no | CNNs with small batches (EM!) |
Real training runs of a small MLP on a Hall–Petch-type dataset (train = blue, validation = orange; always plot both, log scale). LR too high: noisy, no progress. LR too low: smooth but painfully slow. Healthy: fast drop, validation plateaus close to train. Overfitting: training loss → 0 while validation stalls — a large net on \(N=30\) memorises. Figure: img/make_figures.py.

img/make_figures.py.model.eval()).GroupKFold split with the same preprocessing.StandardScaler fit on train only.GroupKFold split, RF/GBM baseline alongside.week06_tiny_mlp.ipynbnotebooks/week06_tiny_mlp.ipynb — CPU only, < 1 min.
Material for questions and self-study — not part of the 90-minute lecture path.
Training loss (blue) and validation R² (orange) versus epoch. Training loss decreases monotonically; validation R² peaks around epoch 70 and then levels off — the optimal stopping point.
StandardScaler (train only) → 4 → 64 → 64 → 1, ReLU, identity output → Adam \(\eta = 10^{-3}\), batch 32, early stopping (n_iter_no_change = 30).GroupKFold split; if the MLP does not win, the data are too small or the relation too simple.
©Philipp Pelz - FAU Erlangen-Nürnberg - Data Science for Electron Microscopy