FAU Erlangen-Nürnberg
Institute of Micro- and Nanostructure Research
notebooks/week09_autoencoder_vae.ipynb — denoising AE vs PCA, anomaly detection, VAE, GMM + BIC, t-SNE sweep, β-sweep. CPU, ≈1 min.Four families: clustering, dimensionality reduction, density estimation, generative models.
Assign \(N\) spectra \(\mathbf{x}_i\in\mathbb{R}^d\) to \(K\) groups \(C_k\) by minimising the within-cluster sum of squares
\[ J = \sum_{k=1}^{K} \sum_{i \in C_k} \|\mathbf{x}_i - \boldsymbol\mu_k\|^2. \]

\[ p(\mathbf{x}) = \sum_{k=1}^{K} \pi_k \,\mathcal{N}(\mathbf{x};\boldsymbol\mu_k,\Sigma_k), \quad \sum_k \pi_k = 1. \]

Hidden label \(c_i\): \(p(\mathbf{x}_i, c_i{=}k)=\pi_k\,\mathcal{N}(\mathbf{x}_i;\boldsymbol\mu_k,\Sigma_k)\). \(\log\sum_k(\cdot)\) → no closed form. EM alternates:
E-step (Bayes’ rule): \[ \gamma_{ik}=\frac{\pi_k\,\mathcal{N}(\mathbf{x}_i;\boldsymbol\mu_k,\Sigma_k)}{\sum_j \pi_j\,\mathcal{N}(\mathbf{x}_i;\boldsymbol\mu_j,\Sigma_j)} \]
M-step (weighted averages), \(N_k=\sum_i\gamma_{ik}\): \[ \boldsymbol\mu_k=\tfrac{1}{N_k}\textstyle\sum_i\gamma_{ik}\mathbf{x}_i,\quad \Sigma_k=\tfrac{1}{N_k}\sum_i\gamma_{ik}(\mathbf{x}_i-\boldsymbol\mu_k)(\mathbf{x}_i-\boldsymbol\mu_k)^{\top},\quad \pi_k=\tfrac{N_k}{N} \]

reg_covar), diagonal \(\Sigma\), or fit in a latent space.Data on a curved low-dimensional surface; PCA fits a flat plane, a nonlinear AE follows the curve.
1024-channel spectrum → 8-D latent code \(\mathbf{z}\) → reconstruction.
\[ \mathcal{L}(\theta, \phi) = \frac{1}{N}\sum_{i=1}^{N} \bigl\|\mathbf{x}_i - g_\theta(f_\phi(\mathbf{x}_i))\bigr\|^2 \qquad\text{— no labels; the input is the target.} \]

Train on corrupted inputs, clean targets Vincent, Pascal et al., (2008): \[\mathcal{L} = \frac{1}{N}\sum_i \|\mathbf{x}_i - g_\theta(f_\phi(\tilde{\mathbf{x}}_i))\|^2\]

AE latent codes of 400 synthetic spectra (4 phases): islands without labels; stars = k-means centroids.
Train on normal spectra only; unseen phase D reconstructs ~25× worse and is flagged (left); error map (right).
Left: AE latent with gaps — sampling there decodes to nonsense. Right: VAE latent, pulled toward \(\mathcal{N}(0,I)\), space-filling.
Decoder (generative): \(\ \mathbf{z}\sim\mathcal{N}(\mathbf{0},I),\quad \mathbf{x}\sim p_\theta(\mathbf{x}\mid\mathbf{z})\)
Encoder (inference): \(\ q_\phi(\mathbf{z}\mid\mathbf{x})=\mathcal{N}\big(\boldsymbol\mu_\phi(\mathbf{x}),\operatorname{diag}\boldsymbol\sigma^2_\phi(\mathbf{x})\big)\)

We want \(\log p_\theta(\mathbf{x})=\log\int p_\theta(\mathbf{x}\mid\mathbf{z})\,p(\mathbf{z})\,d\mathbf{z}\) — intractable for a neural decoder.

For \(q=\mathcal{N}(\boldsymbol\mu,\operatorname{diag}\boldsymbol\sigma^2)\), \(p=\mathcal{N}(\mathbf{0},I)\): \[ \mathrm{KL}(q\,\|\,p)=\tfrac12\sum_{j=1}^{k}\big(\mu_j^2+\sigma_j^2-\log\sigma_j^2-1\big) \]
No Monte Carlo for the KL — only the reconstruction term uses the sampled \(\mathbf{z}\).
mu, logvar = encoder(x_noisy) # predict log σ² → σ > 0 for free
eps = torch.randn_like(mu)
z = mu + torch.exp(0.5 * logvar) * eps # reparameterisation
x_hat = decoder(z)
recon = ((x_hat - x_clean)**2).sum(1) / (2 * SIGMA2)
kl = 0.5 * (mu**2 + logvar.exp() - logvar - 1).sum(1)
loss = (recon + beta * kl).mean() # = −ELBO (β = 1)
loss.backward(); opt.step()
2-D denoising VAE: (a) −ELBO terms, (b) phase latents inside the prior’s 2σ disc, (c) spectra decoded from prior samples.
Recon MSE: noisy 0.00449 · PCA 0.00205 · AE 0.00038 · VAE (\(k\)=2) 0.00017 · ARI 1.00.
Problem: atom-centred STEM patches differ by rotation and sub-pixel shift; a plain VAE spends its latents on orientation.
The same 400 4-D AE latent codes embedded with t-SNE at perplexity 2, 5, 30, 100; ρ = rank correlation of inter-phase distances with the 4-D latent.
Traps
Checklist
Notebook: (a) BIC picks K = 4; (b) GMM ellipses and the latent path of Fe₂O₃/FeO mixtures; (c) responsibilities along the path.
A 50/50 Fe₂O₃/FeO interface pixel is labelled “Fe₃O₄” with \(\gamma = 1.00\). A responsibility is a probability under the fitted mixture, not physical truth — check spatial context, reference spectra, linear unmixing (Week 3).
Synthetic 4D-STEM scan (two grains, bending, thickness gradient) → 3 latent coordinates per probe position. \(z_1\) separates grains; \(z_2\) tracks thickness — a nuisance; GMM recovers the boundary (white).
notebooks/week09_autoencoder_vae.ipynbMaterial moved out of the 90-minute lecture path; not examined beyond what the main slides cover.
| K-means | GMM | |
|---|---|---|
| Assignment | hard (one cluster) | soft (probability vector) |
| Cluster shape | spherical | ellipsoidal (full \(\Sigma\)) |
| Speed | fast | slower (covariance update) |
| Boundary handling | abrupt | smooth |
| Choosing \(K\) | elbow + silhouette | BIC (penalised log-likelihood) |
| High-dim risk | Euclidean dominance | covariance explosion (\(d^2\) params) |
Rule of thumb: k-means for fast exploration; GMM when clusters overlap, vary in shape, or you need probabilities for downstream decisions. In both cases, prefer clustering latent codes over raw spectra Murphy, Kevin P., (2012).
Left: t-SNE (perplexity=30) — 2D embedding of 240 latent codes from a 10-D AE. Four iron-oxide phases form islands. Warning: inter-island distances are not metric — the gap between Fe₂O₃ and FeO in this plot does not reflect their spectral similarity. Right: UMAP (n_neighbors=15) — same codes, tighter clusters, better preservation of global structure. Use UMAP for 2026 pipelines; t-SNE is a useful diagnostic.
| Vanilla AE | VAE | |
|---|---|---|
| Encoder output | a point \(\mathbf{z}\) | a distribution \(\mathcal{N}(\boldsymbol\mu,\operatorname{diag}\boldsymbol\sigma^2)\) |
| Loss | reconstruction | \(-\)ELBO = reconstruction + KL |
| Latent structure | arbitrary scale, holes | ≈ \(\mathcal{N}(\mathbf{0},I)\), space-filling |
| Sampling / interpolation | unreliable | \(\mathbf{z}\sim\mathcal{N}(\mathbf{0},I)\) → decode |
| Per-sample uncertainty | none | \(\boldsymbol\sigma_\phi(\mathbf{x})\) (approximate) |
| Typical EM use | denoising, compression, anomaly scores | smooth latent maps, disentangling (rVAE), generation |
Generative adversarial networks (GANs) — sharper samples, no encoder, no likelihood — return in Week 13 together with diffusion models as priors for inverse problems Goodfellow, Ian et al., (2014).
Workflow of Kalinin, Sergei V et al. (2021) for e-beam-driven dynamics in Si-doped graphene: a deep CNN semantically segments each STEM frame into atom types, sub-images are cropped around atomic positions, and an (r)VAE compresses them into disentangled building blocks (blue: rotation/shift latents; red: content latents). Reproduced from Kalinin et al., Sci. Adv. 7, eabd5084 (2021), open access.
Descriptors for graphene sub-images with small (top) and large (bottom) windows. (A, D) GMM classes: many statistical components, rotation spread over several classes. (B, E) Convolutional AE latent grid: sublattices separate, but rotation, shift and defects are entangled. (C, F) rVAE latent grid: with rotation removed, only one latent direction varies (degree of crystallinity); with larger windows a second direction encodes proximity to edges. From Kalinin, Sergei V et al. (2021).
Complete pipeline from raw EELS/EDS spectra to actionable outputs. Raw spectra enter a denoising AE/VAE encoder; the latent space \(z\) branches into (1) a phase map via k-means/GMM, (2) denoised spectra via the decoder, and (3) anomaly flags where reconstruction error exceeds a threshold. All three outputs require zero labels — and all three need the validation checklist before they become claims.

©Philipp Pelz - FAU Erlangen-Nürnberg - Data Science for Electron Microscopy