Data Science for Electron Microscopy
Week 6: Neural networks & backpropagation

Prof. Dr. Philipp Pelz

FAU Erlangen-Nürnberg

Institute of Micro- and Nanostructure Research

FAU Logo IMN Logo CENEM Logo ERC Logo Eclipse Logo

Recap: where we left off

  • Week 5: image → segmentation → regionprops / local-environment descriptors → trees, RF and GBM, validated with GroupKFold by image.
  • Baseline ladder: mean → linear/Ridge → RF/GBM → neural network. Each rung must beat the one below on the same honest split.
  • Week 4: training = minimising a loss, \(\theta_{t+1} = \theta_t - \eta\,\nabla_\theta\hat{R}(\theta_t)\).
  • Today’s question: what if the feature map \(\boldsymbol\phi(\mathbf{x};\theta)\) itself is learned end-to-end by gradient descent — and how do we get its gradients?

Road map and learning outcomes

Road map (≈90 min)

  1. Hand-crafted vs learned features (Hall–Petch)
  2. Perceptron, XOR, the dense layer
  3. Activations, softmax & cross-entropy, vanishing gradients
  4. Autograd & backprop — worked 2-layer example, gradient check
  5. Initialisation (Xavier/He) & normalisation (batch/layer/group)
  6. Diagnostics: loss curves, LR finder, overfit-one-batch
  7. Regularisation, baseline ladder, when not to use a NN

After today you can …

  • write an MLP forward pass in matrix form and count its parameters;
  • match output activation and loss;
  • backprop by hand through a small network and verify with autograd + finite differences;
  • explain vanishing/exploding gradients and the fixes (ReLU, init, normalisation);
  • diagnose a run from its loss curves;
  • justify a NN only if it beats RF/GBM on the same split.

Self-study: notebooks/week06_tiny_mlp.ipynb

Hall–Petch: hand-crafted vs learned features

Left: linear regression on the hand-crafted \(1/\sqrt{d}\) feature achieves \(R^2\approx0.91\) (225 training samples). Right: a small MLP on raw grain size \(d\) only reaches \(R^2\approx0.65\) — with limited data, domain knowledge still wins over learned features.

  • \(H = H_0 + k\,d^{-1/2}\): the engineer supplies \(\phi(d) = d^{-1/2}\). With many interacting inputs (composition, \(T_\text{proc}\), cooling rate) the right \(\phi\) is unknown — a network learns it from data.

From fixed to learned feature maps

  • Fixed basis (up to Week 5): \(\hat{y} = \sum_j w_j\,\phi_j(\mathbf{x})\) — Fourier, polynomial, \(1/\sqrt{d}\), regionprops; only \(w_j\) is learned.
  • Neural network: \(\phi_j(\mathbf{x};\theta)\) is learnable. The architecture decides which features are easy to learn; training finds them.

Hand-crafted features encode what the engineer knows. Learned features encode what the data contains.

The perceptron: one neuron, one hyperplane

A single perceptron: inputs \(x_i\) weighted by \(w_i\), bias \(b\) added, sum passed through an activation \(\sigma\).
  • \(a = \sigma(\mathbf{w}^\top\mathbf{x} + b)\)
  • Boundary \(z = \mathbf{w}^\top\mathbf{x} + b = 0\) is a hyperplane (a line in 2-D).
  • Trained with gradient descent on a loss, this is Week 4’s linear / logistic regression.
  • One neuron = one straight line. It cannot curve.

XOR: why a hidden layer is needed

AND-like data: one line separates the classes. XOR: no single line can.

An MLP with hidden layers learns a curved boundary that classifies all four XOR quadrants.
  • The hidden layer warps the input space into learned features where a final linear boundary suffices.

The dense layer: matrix view and parameter count

An MLP: hidden layers learn internal representations; the output layer is a linear map on them.

\[\mathbf{z} = W\mathbf{x} + \mathbf{b}, \qquad \mathbf{a} = \sigma(\mathbf{z})\]

  • \(W \in \mathbb{R}^{M \times D}\), \(\mathbf{b} \in \mathbb{R}^M\). Batch: \(A = \sigma(WX + \mathbf{b})\), bias broadcasts.
  • Parameters: \(M(D+1)\). Example: \(D=20\), \(M=64\) → 1,344.

Why non-linearity is non-negotiable

  • Two purely linear layers: \[W^{(2)}\bigl(W^{(1)}\mathbf{x}+\mathbf{b}^{(1)}\bigr)+\mathbf{b}^{(2)} = \underbrace{(W^{(2)}W^{(1)})}_{\tilde{W}}\mathbf{x}+\tilde{\mathbf{b}}\]
  • Any depth of linear layers = one affine map. Depth without nonlinearity is no depth at all — and XOR stays unsolvable.

Activation functions: what to use where

Step, sigmoid, tanh, ReLU, Leaky ReLU. ReLU preserves the gradient for positive inputs; sigmoid and tanh saturate.

  • Hidden layers: ReLU \(\max(0,z)\) — gradient exactly 1 for \(z>0\), cheap. Step: zero gradient, untrainable. Sigmoid \(\sigma' \le 0.25\) and tanh saturate for \(|z| \gtrsim 3\) — used at outputs, not inside.

Output layer: softmax and matched losses

  • \(K\) classes: \(\hat{p}_k = e^{z_k} / \sum_j e^{z_j}\) — a probability vector. Cross-entropy \(\mathcal{L} = -\sum_k y_k \log \hat{p}_k\); the output gradient is simply \(\hat{\mathbf{p}} - \mathbf{y}\).
Task Output activation Loss
Regression identity MSE or MAE
Binary classification sigmoid binary cross-entropy
Multi-class softmax categorical cross-entropy

Vanishing gradients: why ReLU won

Left: sigmoid gradient ≤ 0.25, ReLU gradient = 1 for \(z>0\). Right: at depth 8 sigmoid gradients have shrunk by four orders of magnitude; ReLU gradients stay at full strength.
  • The chain rule multiplies one \(\sigma'\) per layer.
  • Sigmoid: \(\le 0.25^8 \approx 2\times10^{-4}\) after 8 layers → early layers barely learn Goodfellow, Ian et al., (2016).
  • ReLU: factor 1 for active units → depth becomes trainable.

Training a network: the same ERM loop

\[\hat{R}(\theta) = \frac{1}{N}\sum_{i=1}^{N} \mathcal{L}\!\bigl(f_\theta(\mathbf{x}_i),\, y_i\bigr), \qquad \theta \leftarrow \theta - \eta\,\nabla_\theta \hat{R}\]

  • \(\theta\) = all weights and biases of all layers; \(f_\theta\) is a composition of nonlinear maps; the loss surface is non-convex.
  • SGD / Adam as before — Adam at \(\eta = 10^{-3}\) is the usual start.
  • The only new problem: computing \(\nabla_\theta\hat{R}\) for \(10^4\)–\(10^9\) parameters.

Why we need backpropagation

Method Passes needed Cost for \(W\) params Exact?
Finite differences \(W+1\) forward \(O(W^2)\) no (\(\epsilon\) error)
Symbolic — expression swell yes
Backprop (reverse AD) 1 forward + 1 backward \(O(W)\) yes

Chain rule: the only calculus you need

  • Univariate: \(y = f(g(x)) \Rightarrow dy/dx = f'(g(x))\,g'(x)\).
  • Multivariate: \(L\) depends on \(x\) through \(u_1,\dots,u_k\): \[\frac{\partial L}{\partial x} = \sum_{i=1}^{k} \frac{\partial L}{\partial u_i}\,\frac{\partial u_i}{\partial x}\] → multiply along each path, sum over paths.
  • Vector form: Jacobians compose by matrix multiplication.

Autograd: backprop on the computational graph

Computational graph for one neuron: forward pass (black, left to right) computes the loss; backward pass (red dashed, right to left) propagates gradients via the chain rule at each node.
  • Forward: evaluate nodes left to right; store intermediates.
  • Backward: right to left, multiply the incoming gradient by each node’s local derivative.
  • PyTorch/JAX build the graph as your code runs: loss.backward() fills .grad.

Forward mode vs reverse mode

  • Forward mode: one pass per input — efficient when #inputs ≪ #outputs.
  • Reverse mode: one pass per output — efficient when #outputs ≪ #inputs.
  • Training maps \(10^4\)–\(10^9\) parameters → one scalar loss ⇒ reverse mode = backpropagation.
  • Price: store all \(\mathbf{z}^{(\ell)}, \mathbf{a}^{(\ell)}\) — memory, not compute, is the limit for large EM images.

Worked example (1/3): the network and the forward pass

Network: 1 input → 2 ReLU hidden units → 1 linear output, squared loss.

\[ \begin{aligned} \mathbf{z}^{(1)} &= \mathbf{W}^{(1)}x + \mathbf{b}^{(1)}, &\quad \mathbf{a}^{(1)} &= \mathrm{ReLU}(\mathbf{z}^{(1)})\\ \hat{y} &= \mathbf{w}^{(2)\top}\mathbf{a}^{(1)} + b^{(2)}, &\quad L &= \tfrac12(\hat{y}-y)^2 \end{aligned} \]

Numbers: \(x = 1,\; y = 0.5\), \(\mathbf{W}^{(1)} = \begin{pmatrix}2\\-1\end{pmatrix}\), \(\mathbf{b}^{(1)} = \begin{pmatrix}-0.5\\0.5\end{pmatrix}\), \(\mathbf{w}^{(2)} = \begin{pmatrix}1\\2\end{pmatrix}\), \(b^{(2)} = 0\).

  • \(\mathbf{z}^{(1)} = (2\cdot1-0.5,\; -1\cdot1+0.5) = (1.5,\,-0.5)\)
  • \(\mathbf{a}^{(1)} = \mathrm{ReLU}(\mathbf{z}^{(1)}) = (1.5,\; 0)\) ← unit 2 is off
  • \(\hat{y} = 1\cdot1.5 + 2\cdot0 + 0 = 1.5\)
  • \(L = \tfrac12(1.5-0.5)^2 = 0.5\)
  • Store \(\mathbf{z}^{(1)}, \mathbf{a}^{(1)}, \hat y\) — the backward pass needs them.

Worked example (2/3): the backward pass

  • Output: \(\delta^{(2)} = \dfrac{\partial L}{\partial \hat y} = \hat{y} - y = 1.0\)
  • Output-layer parameters: \(\dfrac{\partial L}{\partial \mathbf{w}^{(2)}} = \delta^{(2)}\,\mathbf{a}^{(1)} = (1.5,\; 0)\), \(\quad\dfrac{\partial L}{\partial b^{(2)}} = \delta^{(2)} = 1.0\)
  • Into the hidden layer: \(\dfrac{\partial L}{\partial \mathbf{a}^{(1)}} = \mathbf{w}^{(2)}\,\delta^{(2)} = (1.0,\; 2.0)\)
  • Through ReLU: \(\boldsymbol{\delta}^{(1)} = \dfrac{\partial L}{\partial \mathbf{z}^{(1)}} = (1.0,\,2.0) \odot \mathrm{ReLU}'(\mathbf{z}^{(1)}) = (1.0,\,2.0)\odot(1,\,0) = (1.0,\; 0)\)
  • Hidden-layer parameters: \(\dfrac{\partial L}{\partial \mathbf{W}^{(1)}} = \boldsymbol{\delta}^{(1)}\,x = (1.0,\; 0)\), \(\quad \dfrac{\partial L}{\partial \mathbf{b}^{(1)}} = \boldsymbol{\delta}^{(1)} = (1.0,\; 0)\)
  • GD step (\(\eta = 0.1\)): \(w^{(2)}_1: 1 \to 0.85\), \(W^{(1)}_1: 2 \to 1.9\), \(b^{(1)}_1: -0.5\to-0.6\); unit 2’s incoming weights do not move this step.

Worked example (3/3): the whole thing on one graph

Forward values (black) and backward gradients (red) for the worked example. Each weight gradient = incoming error signal × stored forward input. The inactive ReLU unit (\(z_2^{(1)}<0\)) passes no gradient. Figure: img/make_figures.py.

The delta recursion and the backprop algorithm

  1. Forward: \(\mathbf{z}^{(\ell)} = \mathbf{W}^{(\ell)}\mathbf{a}^{(\ell-1)} + \mathbf{b}^{(\ell)}\), \(\mathbf{a}^{(\ell)} = \sigma(\mathbf{z}^{(\ell)})\); store both (\(\mathbf{a}^{(0)} = \mathbf{x}\)).
  2. Output delta: \(\boldsymbol{\delta}^{(L)} = \nabla_{\mathbf{a}^{(L)}}\mathcal{L} \odot \sigma'(\mathbf{z}^{(L)})\)
  3. Backward, \(\ell = L-1..1\): \[\boldsymbol{\delta}^{(\ell)} = \sigma'(\mathbf{z}^{(\ell)}) \odot \bigl(\mathbf{W}^{(\ell+1)\top}\boldsymbol{\delta}^{(\ell+1)}\bigr)\]
  4. Gradients: \(\nabla_{\mathbf{W}^{(\ell)}}\mathcal{L} = \boldsymbol{\delta}^{(\ell)}\mathbf{a}^{(\ell-1)\top}\), \(\;\nabla_{\mathbf{b}^{(\ell)}}\mathcal{L} = \boldsymbol{\delta}^{(\ell)}\)
  • \(\mathbf{W}^{\top}\boldsymbol\delta\) = “sum over paths”.
  • \(\odot\,\sigma'\) = local gate (ReLU passes/blocks, sigmoid shrinks ≤ 0.25).
  • Softmax + CE: \(\boldsymbol\delta^{(L)} = \hat{\mathbf{p}} - \mathbf{y}\).

Autograd in PyTorch: the training step

# the training step of every network in this course
opt.zero_grad()              # gradients ACCUMULATE by default -> reset them
loss = loss_fn(model(xb), yb)
loss.backward()              # reverse-mode AD: fills p.grad for every parameter p
opt.step()                   # theta <- theta - eta * grad  (or Adam)
  • torch.no_grad() for evaluation · model.eval() switches dropout / batch norm to inference mode.

Gradient checking: trust, but verify

  • Central differences per parameter: \(g_j^\text{num} = \dfrac{\mathcal{L}(\theta + \epsilon\mathbf{e}_j) - \mathcal{L}(\theta - \epsilon\mathbf{e}_j)}{2\epsilon}\), error \(O(\epsilon^2)\).
  • Compare with the analytic/autograd gradient via the relative error \[r = \frac{\|\mathbf{g}^\text{auto} - \mathbf{g}^\text{num}\|}{\|\mathbf{g}^\text{auto}\| + \|\mathbf{g}^\text{num}\|}\]
  • Rules of thumb (float64, \(\epsilon \approx 10^{-6}\)): \(r < 10^{-7}\) fine · \(10^{-4}\) suspicious · \(> 10^{-2}\) bug.
  • Use float64, a tiny model, a few samples, and turn off dropout (randomness breaks the check). Avoid kinks: ReLU at \(z \approx 0\) can give spurious mismatches.
  • Notebook: manual NumPy backprop vs loss.backward() vs finite differences on the worked example and a random 2-layer net.

Exploding gradients and gradient clipping

  • Backward pass = product of \(L\) factors \(\mathrm{diag}(\sigma'(\mathbf{z}^{(\ell)}))\,\mathbf{W}^{(\ell+1)\top}\).
  • Factors \(< 1\) → vanish (sigmoid, tiny weights); factors \(> 1\) → explode (large weights) → loss jumps, then NaN.
  • Clipping (safety net): if \(\|\nabla\| > \tau\), rescale to norm \(\tau\) — clip_grad_norm_(model.parameters(), 1.0).
  • Structural fixes keep each factor ≈ 1: ReLU, initialisation, normalisation, residual connections (Week 7).

Initialisation: symmetry breaking and scale

  • Zero / constant init: all hidden units compute the same output, get the same gradient → the layer behaves like one neuron forever. Always wrong (biases may start at 0).
  • Scale: too small → signals shrink layer by layer; too large → explode (ReLU) or saturate (tanh).
  • Xavier / Glorot Glorot, Xavier et al., (2010): \(\mathrm{Var}(w) = \dfrac{2}{D_\text{in}+D_\text{out}}\) — tanh/sigmoid.
  • He / Kaiming He, Kaiming et al., (2015): \(\mathrm{Var}(w) = \dfrac{2}{D_\text{in}}\) — ReLU (a variance; std \(= \sqrt{2/D_\text{in}}\)).

Why \(\mathrm{Var}(w) = 2/D_\text{in}\): the variance argument

  • \(z = \sum_{j=1}^{D_\text{in}} w_j a_j\), independent zero-mean weights: \[\mathrm{Var}(z) = D_\text{in}\,\mathrm{Var}(w)\,\mathbb{E}[a^2]\]
  • Linear/tanh: \(\mathbb{E}[a^2] \approx \mathrm{Var}(z_\text{prev})\) → \(\mathrm{Var}(w) = 1/D_\text{in}\) (Xavier averages in/out).
  • ReLU zeroes half: \(\mathbb{E}[a^2] = \tfrac12\mathrm{Var}(z_\text{prev})\) → \(\mathrm{Var}(w) = 2/D_\text{in}\) (He).
  • Otherwise a factor \(c^L\) after \(L\) layers: \(0.5^{20} \approx 10^{-6}\), \(2^{20} \approx 10^{6}\).

Initialisation in action: activation scale through depth

Standard deviation of activations layer by layer in a 20-layer, width-256 network fed with unit-variance input. ReLU (left): small weights → activations vanish, \(\mathcal{N}(0,1)\) → explode, Xavier → slow decay, He → constant. tanh (right): large weights saturate at ±1, small weights vanish, Xavier/He stay usable. Figure: img/make_figures.py.

Normalisation: batch, layer and group norm

  • Inputs, always: StandardScaler fit on the training split only (Week 4 leakage rule).
  • BatchNorm Ioffe, Sergey et al., (2015): per unit, over the mini-batch: \(\tilde{z}_i = \gamma\,\dfrac{z_i - \mu_\mathcal{B}}{\sqrt{\sigma^2_\mathcal{B} + \epsilon}} + \beta\). Running statistics at inference → model.eval()!
  • LayerNorm Ba, Jimmy Lei et al., (2016) / GroupNorm Wu, Yuxin et al., (2018): normalise within one sample → batch-size independent.
averages over batch-size dependent train ≠ eval typical use
BatchNorm batch (per feature/channel) yes yes CNNs, large batches
LayerNorm features of one sample no no transformers (W10), MLPs
GroupNorm channel groups of one sample no no CNNs with small batches (EM!)

Reading loss curves: four patterns you must recognise

Real training runs of a small MLP on a Hall–Petch-type dataset (train = blue, validation = orange; always plot both, log scale). LR too high: noisy, no progress. LR too low: smooth but painfully slow. Healthy: fast drop, validation plateaus close to train. Overfitting: training loss → 0 while validation stalls — a large net on \(N=30\) memorises. Figure: img/make_figures.py.

Learning-rate range test (“LR finder”)

LR range test on a small MLP: η increases exponentially every mini-batch; the loss is flat for tiny η, drops steeply in a useful window, and diverges when η is too large. Figure: img/make_figures.py.
  • Multiply η by a constant every mini-batch, record the loss, stop at blow-up Smith, Leslie N., (2017).
  • Pick η in the steepest region, 3–10× below the minimum.
  • One short pass — far cheaper than a grid search.

Sanity checks before a long run

  1. Initial loss: standardised regression MSE ≈ 1; \(K\)-class softmax \(\mathcal{L}_0 \approx \ln K\) (1.39 for 4 phases).
  2. Overfit one batch: ~16 samples, no regularisation → training loss must reach ≈ 0. If not, it is a bug, not lack of data.
  3. Gradient check custom losses/layers.
  4. Per-layer gradient norms: early ≪ late → vanishing; growing → exploding.
  5. Look at predictions: parity plots, masks, worst errors.

Overfitting: early stopping and dropout

  • A 2 × 64 MLP has > 5000 parameters — often more than the EM samples you have.
  • Cures, in order: more data · early stopping (keep the best-validation epoch) · dropout · smaller network · weight decay.
  • Dropout Srivastava, Nitish et al., (2014): zero a fraction \(p \approx 0.2\)–\(0.5\) of hidden units per training pass → no co-adaptation, an implicit ensemble. Off at test time (model.eval()).

The baseline ladder: a NN must earn its place

  • mean → linear/Ridge → RF/GBM → NN, every rung on the same GroupKFold split with the same preprocessing.
  • Tabular descriptors at small/medium \(N\): tree ensembles are very hard to beat Grinsztajn, Léo et al., (2022).
  • Raw images/spectra: the NN (with the right inductive bias, Week 7) is the only rung that learns features itself.
  • Notebook (Hall–Petch, test \(R^2\)): mean ≈ 0 · Ridge raw 0.80 · MLP 0.65 · RF 0.86 · GBM 0.88 · Ridge on \(1/\sqrt{d}\) 0.91.

When NOT to use a deep net for EM data

  • Small tabular data (\(N < 200\) specimens): Ridge or GBM almost always generalises better.
  • Known physics (\(H \propto d^{-1/2}\)): embed it rather than learn it.
  • Mandatory interpretability: white-box models (Hall–Petch, CALPHAD) when the mechanism is known.
  • Images as flat vectors: a dense layer on 1024×1024 pixels needs \(\sim10^9\) weights — wrong architecture → CNN (Week 7).
  • Universal approximation says a wide net can represent any continuous function — not that GD finds it, that your data suffice, or that it survives distribution shift.

The complete training recipe

  1. Architecture: 2–3 hidden layers, 32–128 units, ReLU.
  2. Initialise: He for ReLU (framework default), never zero.
  3. Normalise inputs: StandardScaler fit on train only.
  4. Sanity checks: initial loss, overfit one batch, LR range test.
  5. Train: Adam \(\eta = 10^{-3}\), batch 32–64, early stopping on validation loss.
  6. Evaluate honestly: same GroupKFold split, RF/GBM baseline alongside.

Self-study notebook: week06_tiny_mlp.ipynb

  • Notebook: notebooks/week06_tiny_mlp.ipynb — CPU only, < 1 min.
    • Ridge on \(1/\sqrt{d}\) vs MLP on raw \((d, T)\), plus RF/GBM on the same split.
    • Manual backprop vs autograd vs central differences on the worked example (relative error < \(10^{-8}\)).
    • Exercise: width and activation (ReLU vs tanh).
  • Open In Colab — complete before Week 7.

Summary: the week in six points

  • Neuron \(\sigma(\mathbf{w}^\top\mathbf{x} + b)\); MLP = composed layers; nonlinearity makes depth meaningful. ReLU inside; matched output/loss.
  • Backprop = reverse-mode autodiff: store activations, run \(\boldsymbol\delta^{(\ell)} = \sigma'(\mathbf{z}^{(\ell)})\odot\mathbf{W}^{(\ell+1)\top}\boldsymbol\delta^{(\ell+1)}\); weight gradient \(\boldsymbol\delta\,\mathbf{a}^\top\); cost ≈ one forward pass.
  • Verify with central finite differences (float64, relative error).
  • Trainability: ReLU, He/Xavier init, batch/layer norm keep per-layer factors ≈ 1; clip as a safety net.
  • Diagnostics: train + val loss curves, LR range test, initial loss, overfit one batch.
  • Honesty: same GroupKFold split, RF/GBM baseline; a NN must earn its place.

Must-know for the exam

  1. Dense layer \(\mathbf{a} = \sigma(\mathbf{W}\mathbf{x} + \mathbf{b})\), parameter count \(M(D+1)\); stacked linear layers collapse to one affine map.
  2. Output/loss pairs: identity + MSE, sigmoid + BCE, softmax + CE (gradient \(\hat{\mathbf{p}} - \mathbf{y}\)).
  3. Hand-compute forward and backward pass of a 2-layer ReLU network (worked example).
  4. Delta recursion and \(\nabla_{\mathbf{W}^{(\ell)}} = \boldsymbol\delta^{(\ell)}\mathbf{a}^{(\ell-1)\top}\); backprop costs \(O(W)\), finite differences \(O(W^2)\).
  5. Gradient check: central differences, relative error, float64.
  6. Vanishing (\(\sigma' \le 0.25\)) / exploding gradients; ReLU, clipping.
  7. Zero init fails (symmetry); Xavier \(2/(D_\text{in}+D_\text{out})\), He \(2/D_\text{in}\) — and why.
  8. BatchNorm vs LayerNorm: what is averaged, train vs eval, small-batch caveat.
  9. Read loss curves; LR range test; initial loss ≈ ln K; overfit-one-batch.
  10. Baseline ladder: NN vs RF/GBM on the same GroupKFold split.

Next week (Week 7): CNNs & U-Nets for microscopy

  • Problem: a flattened 1024×1024 image into a dense layer needs \(\sim 10^9\) weights — with \(N < 1000\) images the model memorises noise.
  • Fix: build in locality and weight sharing → convolution, a sparse, weight-tied matrix multiply.
  • Week 7: receptive fields, pooling; VGG → ResNet (skip connections + batch norm — today’s tools); U-Net segmentation with Dice/IoU; saliency and Grad-CAM.
  • Inductive-bias arc: MLP (no structure) → CNN (grid, W7) → GNN / transformer (graphs and sets, W10).

Continue

Backup slides

Material for questions and self-study — not part of the 90-minute lecture path.

The information bottleneck: from micrograph to scalar

  • A 1024×1024 BSE micrograph has \(\sim 10^6\) pixels; an ASTM grain-size number discards \(\sim 10^6{:}1\).
  • A hand-crafted descriptor answers exactly the question it was designed for, and nothing else Sandfeld, Stefan et al., (2024).
  • Learned representations let the network decide which aspects of the field matter for the target — adaptive, not blind compression — given enough data and honest validation.

ADALINE and gradient descent on a neuron

  • History: MCP neuron (1943, fixed weights) → Perceptron (1957, error-correction rule, hard threshold) → ADALINE (1960): continuous output, MSE, gradient descent: \[\nabla_\mathbf{w}\hat{R} = \frac{2}{N}\mathbf{X}^T(\mathbf{X}\mathbf{w} - \mathbf{y})\]
  • This is exactly the Week 4 linear-regression GD update, at the level of a single neuron.

MLPs learn hierarchical features

  • Layer 1: simple combinations of raw features (“large grain AND high temperature”).
  • Layer 2: combinations of those (“… AND low Cr fraction”). Output: linear map on the last hidden layer.
  • Universal approximation: one wide hidden layer can approximate any continuous function Goodfellow, Ian et al., (2016); in practice depth is more parameter-efficient than width.

Historical activations: step, sigmoid, tanh

  • Step \(\mathbb{1}[z \ge 0]\): derivative zero everywhere — no gradient learning (MCP neuron, original perceptron).
  • Sigmoid \((1+e^{-z})^{-1}\): \(\sigma' = \sigma(1-\sigma) \le 0.25\), saturates for \(|z|>3\); dominated 1986–2012; today only at binary outputs.
  • Tanh: zero-centred, \(\tanh' \le 1\), still saturates. ReLU (AlexNet, 2012): gradient exactly 1 for \(z>0\) — made 10–100-layer networks trainable.

Training loss and early stopping

Training loss (blue) and validation R² (orange) versus epoch. Training loss decreases monotonically; validation R² peaks around epoch 70 and then levels off — the optimal stopping point.

MLP for materials property prediction: a complete example

  • Problem: Vickers hardness from grain size, processing temperature, Cr and Mn fraction.
  • Pipeline: StandardScaler (train only) → 4 → 64 → 64 → 1, ReLU, identity output → Adam \(\eta = 10^{-3}\), batch 32, early stopping (n_iter_no_change = 30).
  • Baseline: Ridge and RF/GBM on the same features and GroupKFold split; if the MLP does not win, the data are too small or the relation too simple.

References

Materials data science, Stefan Sandfeld & others.
Deep learning, Ian Goodfellow, Yoshua Bengio, & Aaron Courville.
Learning representations by back-propagating errors, Nature, David E. Rumelhart, Geoffrey E. Hinton, & Ronald J. Williams.
Automatic differentiation in machine learning: A survey, Journal of Machine Learning Research, Atılım Güneş Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul, & Jeffrey Mark Siskind.
Pattern recognition and machine learning, Christopher M. Bishop.
Understanding the difficulty of training deep feedforward neural networks, International conference on artificial intelligence and statistics (AISTATS), Xavier Glorot & Yoshua Bengio.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification, International conference on computer vision (ICCV), Kaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun.
Batch normalization: Accelerating deep network training by reducing internal covariate shift, International conference on machine learning (ICML), Sergey Ioffe & Christian Szegedy.
Layer normalization, arXiv preprint arXiv:1607.06450, Jimmy Lei Ba, Jamie Ryan Kiros, & Geoffrey E. Hinton.
Group normalization, European conference on computer vision (ECCV), Yuxin Wu & Kaiming He.
Cyclical learning rates for training neural networks, 2017 IEEE winter conference on applications of computer vision (WACV), Leslie N. Smith.
Dropout: A simple way to prevent neural networks from overfitting, Journal of Machine Learning Research, Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, & Ruslan Salakhutdinov.
Why do tree-based models still outperform deep learning on tabular data?, Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, Léo Grinsztajn, Edouard Oyallon, & Gaël Varoquaux.